ML.NET vs PyTorch for inference: CPU and GPU comparison

ML.NET and PyTorch inference paths
ML.NET and PyTorch inference paths

ML.NET and PyTorch fit different parts of the machine learning lifecycle. For a .NET team serving a PyTorch-trained model, the real choice is usually between keeping a PyTorch service and exporting to ONNX for in-process inference. The historical ResNet18 test below is useful as a benchmarking lesson, but its timing flaws mean the reported numbers cannot rank either framework today.

ML.NET vs PyTorch: the short answer

PyTorch is usually the stronger choice for developing and training custom deep learning models. ML.NET is designed for machine learning in .NET applications, including classical machine learning, selected deep learning training APIs, and inference with imported TensorFlow or Open Neural Network Exchange (ONNX) models.

For a model trained in PyTorch, a .NET team has three practical deployment choices:

  1. Keep the model and its preprocessing in a Python or C++ PyTorch runtime behind a service boundary.
  2. Export the model to ONNX and score it inside the .NET process through ML.NET, which uses ONNX Runtime for ONNX inference.
  3. Use the ONNX Runtime C# API directly when the application needs ONNX scoring without ML.NET's IDataView, transforms, or training APIs.

That makes ML.NET vs PyTorch an architecture decision more than a direct runtime contest.

CriterionPyTorchML.NET
Best fitCustom deep learning, research, training, and Python-native inference.NET-native machine learning pipelines and application integration
TrainingBroad control over neural network architecture and training loopsClassical ML plus selected deep learning trainers powered by components such as TorchSharp and TensorFlow.NET
Imported modelsNative PyTorch artifacts and export pathsTensorFlow and ONNX models inside .NET applications
GPU pathPyTorch device backends and compilation toolingHardware acceleration available for supported trainers and imported-model runtimes
Deployment shapePython service, C++ runtime, exported program, or ONNXIn-process .NET pipeline, often with ONNX Runtime for a PyTorch-trained model
Main constraintAdds a separate runtime when the rest of the system is .NETONNX export coverage, preprocessing parity, and execution-provider support can limit portability

ML.NET does more than ONNX scoring. Its deep learning overview documents custom training APIs as well as imported-model inference. For this ResNet18 comparison, however, the ML.NET path is specifically an ONNX Runtime path.

What the original ResNet18 experiment tested

The original 2021 experiment by Valerii Boldakov ran 10,000 single-item inferences with a pretrained ResNet18 model. Both paths used the same computer:

  • AMD Ryzen 5 3600 CPU
  • NVIDIA GeForce GTX 1660 GPU
  • Ubuntu 18.04

The software stacks differed:

PathHistorical software stack
PyTorchPyTorch 1.7.1, torchvision 0.8.2, CUDA 10.1, CUDA Deep Neural Network library (cuDNN) 7
ML.NET with ONNX RuntimeML.NET 1.5.4, ONNX Runtime GPU 1.6.0, OnnxTransformer 1.5.4, CUDA 10.2, cuDNN 8

The test used random tensors with shape 1 x 3 x 256 x 260. It did not use decoded, resized, normalized images. That distinction matters because the measurement covered model calls rather than a complete image-classification request.

Original PyTorch timing code

The screenshot-only code from the original post is reproduced as text below. It shows the experiment as it ran, including the timing issue discussed later.

resnet_gpu_model = torch.jit.trace( torchvision.models.resnet18(pretrained=True).eval().cuda(), torch.randn(1, 3, 256, 260).cuda(), ) resnet_cpu_model = torch.jit.trace( torchvision.models.resnet18(pretrained=True).eval().cpu(), torch.randn(1, 3, 256, 260), ) gpu_inference_time = [] cpu_inference_time = [] for i in range(10000): print(i) input_tensor = torch.randn(3, 256, 260) input_batch = input_tensor.unsqueeze(0) gpu_input_batch = input_batch.cuda() with torch.no_grad(): start = timer() gpu_output = resnet_gpu_model(gpu_input_batch) end = timer() gpu_inference_time.append(end - start) start = timer() cpu_output = resnet_cpu_model(input_batch) end = timer() cpu_inference_time.append(end - start)

The original screenshot did not include its imports or the definition of timer.

The model was then exported to ONNX with this code:

torch.onnx.export( resnet_cpu_model, input_batch, "resnet18.onnx", export_params=True, do_constant_folding=True, input_names=["input"], output_names=["output"], dynamic_axes={ "input": {0: "batch_size"}, "output": {0: "batch_size"}, }, )

Original ML.NET model and timing code

The C# path defined the model's input and output columns, then created separate GPU and CPU prediction engines:

public class ImageLabels { [ColumnName("output")] public float[] Labels { get; set; } } public class PixelValues { public const int ChannelAmount = 3; public const int ImageWidth = 256; public const int ImageHeight = 260; [VectorType(3, 256, 260)] [ColumnName("input")] public float[] Values { get; set; } } var mlContext = new MLContext(); var gpuModelEstimator = mlContext.Transforms.ApplyOnnxModel( outputColumnName: "output", inputColumnName: "input", modelFile: OnnxModelPath, gpuDeviceId: 0); var cpuModelEstimator = mlContext.Transforms.ApplyOnnxModel( outputColumnName: "output", inputColumnName: "input", modelFile: OnnxModelPath); var data = mlContext.Data.LoadFromEnumerable(new List<PixelValues>()); var gpuModel = gpuModelEstimator.Fit(data); var cpuModel = cpuModelEstimator.Fit(data); var gpuPredictionEngine = mlContext.Model.CreatePredictionEngine<PixelValues, ImageLabels>(gpuModel); var cpuPredictionEngine = mlContext.Model.CreatePredictionEngine<PixelValues, ImageLabels>(cpuModel);

The original loop reused one Stopwatch for both calls:

var randNum = new Random(); var gpuInferenceTime = new List<double>(); var cpuInferenceTime = new List<double>(); for (var i = 0; i < 10000; ++i) { Console.WriteLine(i); var value = new PixelValues { Values = Enumerable .Repeat(0, PixelValues.ChannelAmount * PixelValues.ImageHeight * PixelValues.ImageWidth) .Select(_ => (float)randNum.NextDouble()) .ToArray() }; var sw = new Stopwatch(); sw.Start(); var labels = gpuPredictionEngine.Predict(value).Labels; sw.Stop(); gpuInferenceTime.Add(sw.Elapsed.TotalSeconds); sw.Start(); labels = cpuPredictionEngine.Predict(value).Labels; sw.Stop(); cpuInferenceTime.Add(sw.Elapsed.TotalSeconds); }

Historical reported results

The experiment reported these means and standard deviations:

Historical pathReported mean latencyReported standard deviation
PyTorch CPU26 ms3 ms
PyTorch GPU1 ms0.1 ms
ML.NET with ONNX Runtime CPU16 ms2.9 ms
ML.NET with ONNX Runtime GPU5 ms0.7 ms

These figures are preserved as a historical observation on one machine. They were not rerun for this update and do not establish a current performance ranking.

Why the historical numbers cannot name a winner

Two code-level errors affect the headline comparisons.

First, CUDA work is asynchronous. The PyTorch loop starts and stops a host wall-clock timer around the GPU call without synchronizing the CUDA device. It can therefore measure kernel dispatch instead of completed inference. The official PyTorch benchmark guide calls out CUDA synchronization, warm-up, and thread settings as common sources of benchmark errors.

Second, the C# loop calls sw.Start() before the CPU prediction without calling Reset() or Restart(). A stopped Stopwatch retains its elapsed time, so the saved CPU value includes the preceding GPU interval.

Other limitations also prevent a general comparison:

  • The two paths used different CUDA and cuDNN versions.
  • There was no documented warm-up phase.
  • The input size was 256 x 260, while the usual ResNet18 image pipeline uses 224 x 224 crops.
  • Unnormalized random inputs omitted realistic decoding, resizing, normalization, and postprocessing.
  • The test did not check output parity between PyTorch and the exported ONNX model.
  • CPU thread counts, affinity, power state, and background load were not controlled.
  • The GPU test did not document host-to-device transfers or ONNX Runtime node placement.
  • Only single-request latency was recorded. Throughput, batching, queueing, and concurrency were outside the test.

The original text also described 1 ms versus 26 ms as roughly 10 times faster. The arithmetic is 26 times lower latency. The missing CUDA synchronization still makes that ratio unsuitable for comparison.

How to benchmark PyTorch and ML.NET fairly

A useful benchmark starts with the deployment decision, then measures the same contract on both paths.

  1. Define the timing boundary. Report model-only latency and end-to-end latency separately. End-to-end timing should include serialization, preprocessing, device transfers, postprocessing, and any process or network boundary that production requests cross.
  2. Freeze one model and validate outputs. Use the same weights, precision, input shapes, batch sizes, and preprocessing. Compare PyTorch and ONNX outputs across a representative corpus before comparing speed.
  3. Pin the complete environment. Record the CPU, GPU, driver, operating system, framework, ONNX Runtime, CUDA, cuDNN, execution provider, precision, and model opset. CUDA compatibility is version-specific, so use the applicable compatibility matrix instead of copying the old CUDA 10.x setup.
  4. Warm up each path. Run enough untimed requests to initialize libraries, compile graphs where applicable, populate caches, and reach a stable hardware state.
  5. Synchronize accelerated work. For PyTorch CUDA timing, synchronize before starting and after completing the measured call, or use CUDA events or torch.utils.benchmark correctly.
  6. Control CPU parallelism. Pin intra-operation and inter-operation thread settings. Use the same concurrency target that the service will face.
  7. Verify GPU execution. ONNX Runtime execution providers assign supported nodes or subgraphs to hardware. Selecting the CUDA provider alone does not prove every node ran on the GPU. Inspect provider assignment and CPU fallback. The execution-provider model explains this split.
  8. Account for data movement. If inputs and outputs remain on the GPU, test ONNX Runtime I/O binding. Its I/O binding guidance explains how implicit CPU-to-device and device-to-CPU copies can dominate a run.
  9. Report distributions and capacity. Include median, 95th-percentile (p95), and 99th-percentile (p99) latency, plus throughput at realistic batch sizes and concurrency. Retain cold-start measurements as a separate result.

The core timer logic should look like this, with representative inputs prepared outside the model-only timing boundary:

model.eval() with torch.inference_mode(): for sample in warmup_inputs: _ = model(sample) if device.type == "cuda": torch.cuda.synchronize() latencies_ns = [] for sample in benchmark_inputs: if device.type == "cuda": torch.cuda.synchronize() start = time.perf_counter_ns() _ = model(sample) if device.type == "cuda": torch.cuda.synchronize() latencies_ns.append(time.perf_counter_ns() - start)

In C#, Stopwatch.Restart() prevents the accumulation bug from the historical loop:

foreach (var sample in warmupInputs) _ = predictionEngine.Predict(sample); var latencies = new List<TimeSpan>(); var stopwatch = new Stopwatch(); foreach (var sample in benchmarkInputs) { stopwatch.Restart(); _ = predictionEngine.Predict(sample); stopwatch.Stop(); latencies.Add(stopwatch.Elapsed); }

This C# sample is a single-thread latency skeleton. PredictionEngine is not thread-safe, so a production load test needs a deliberate concurrency design rather than one shared instance. Microsoft documents the constraint in its ASP.NET deployment guidance.

Current deployment paths

Keep inference in PyTorch

Use a PyTorch-native runtime when the model relies on unsupported ONNX operators, custom extensions, Python preprocessing, dynamic control flow, or frequent architecture changes. Set the model to evaluation mode and use torch.inference_mode() where its stricter autograd behavior fits. PyTorch's inference mode documentation notes that inference mode does not call model.eval() for you.

PyTorch also provides torch.compile, export tooling, and ahead-of-time compilation through AOTInductor. Treat each as a deployment option to validate against your model and platform requirements. Do not use the original article's TorchServe recommendation as the default path. The project is no longer actively maintained, and its repository warns that planned updates and security patches have stopped.

Train in PyTorch and run ONNX in .NET

This is often the cleanest mixed-stack design when a stable PyTorch model exports successfully and the .NET application benefits from in-process inference.

  1. Put the PyTorch model in evaluation mode.
  2. Export with representative inputs and an ONNX operator-set (opset) version supported by the target runtime.
  3. Declare dynamic shapes only where the production contract needs them.
  4. Validate output names, shapes, data types, values, and task-level predictions across a fixed corpus.
  5. Load the ONNX model through ML.NET's ApplyOnnxModel, or use ONNX Runtime's C# API directly.
  6. Select the execution provider deliberately, inspect fallback, and benchmark device transfers.
  7. Re-run parity and performance tests whenever the model, exporter, runtime, driver, or hardware changes.

The current PyTorch exporter uses the torch.export-based path when dynamo=True, which is also the documented default. A minimal export is accessible as ordinary code:

model.eval() onnx_program = torch.onnx.export( model, (example_input,), input_names=["input"], output_names=["output"], dynamo=True, verify=True, ) onnx_program.save("resnet18.onnx")

The current ONNX exporter also supports explicit opsets, dynamic shapes, custom translations, export reports, and profiling. Export success is only the first gate. Unsupported or custom operators, shape constraints, numerical differences, and preprocessing mismatches still require tests.

Use ML.NET for a .NET-native model

Choose ML.NET directly when its trainers and transforms fit the problem, especially for classical machine learning inside a .NET application. This route keeps training and inference code in C# or F# and avoids an export boundary.

For a custom neural network already developed in PyTorch, rebuilding it as an ML.NET training pipeline usually adds work without improving the deployment decision. Exporting to ONNX or keeping a PyTorch service preserves a clearer ownership boundary.

Decision guide for .NET teams

SituationRecommended starting pathMain validation
Custom deep learning and Python-heavy preprocessingPyTorch serviceService latency, scaling, and operational ownership
Stable PyTorch model in a .NET applicationPyTorch to ONNX to ML.NETExport coverage, output parity, and end-to-end latency
Stable ONNX model with no need for ML.NET transformsONNX Runtime C# APISession concurrency, provider assignment, and data movement
Classification, regression, forecasting, or recommendation using ML.NET trainersML.NET-native pipelineModel quality and production concurrency
Custom operators or failed ONNX exportKeep the PyTorch runtimePlatform fit and support for the model's operators
Model compute must scale independently from the applicationSeparate inference serviceNetwork overhead, batching, autoscaling, and failure isolation

Choose the smallest architecture that preserves model correctness, operational clarity, and acceptable latency. A single-process .NET deployment reduces service boundaries, while a separate PyTorch service lets the model team move and scale independently.

FAQ

Is ML.NET faster than PyTorch?

There is no framework-wide answer. Performance depends on the model, graph, runtime, execution provider, hardware, precision, input shape, batch size, threading, data transfers, and timing boundary. The 2021 numbers above contain timing errors and cannot establish a winner.

Can ML.NET run a PyTorch model?

ML.NET can consume a compatible ONNX model exported from PyTorch. ApplyOnnxModel does not load a PyTorch checkpoint such as a .pt or .pth file directly. The export must preserve the required operators, shapes, preprocessing, postprocessing, and output behavior.

How does ML.NET vs Python differ from ML.NET vs PyTorch?

Python is a programming language and ecosystem. PyTorch is a machine learning framework whose primary interface is Python. A .NET application can call a Python inference service, load an exported ONNX model through ML.NET, use ONNX Runtime directly, or use another .NET binding. The deployment boundary matters more than the language label.

Is TorchServe still maintained?

No. The TorchServe repository says the project is no longer actively maintained and has no planned updates, fixes, features, or security patches. Existing releases remain available, but new deployments need another serving plan.

Run a decision benchmark

Take one production model, one representative request corpus, and one target environment. Validate PyTorch-to-ONNX parity first, then benchmark the PyTorch service and .NET deployment under the same latency and concurrency requirements. That result will answer the architecture question more reliably than a generic framework ranking.

Related Posts

We use cookies for functional and analytical purposes. Please refer to our Privacy Policy for details.