Edge and On-device AI

Core ML

Core ML is Apple's on-device model format, which automatically routes each part of a model to the CPU, GPU or Neural Engine — and only runs on Apple hardware.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. What Core ML does that is genuinely clever
  5. Getting a model into Core ML
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Core ML is Apple's own format for running models on iPhones, iPads and Macs.

The analogy you have already lived

Look at the back of your charger. An Indian three-pin plug does not fit a British socket. Neither is wrong. They were designed for different walls.

Core ML is Apple's socket. A model in Core ML format fits an iPhone perfectly and draws the best power the device can give. That same file fits nothing else at all.

This is a real trade, not a complaint. You get the best performance available on that hardware, and you get it nowhere else.

Why it exists

Apple designs the chip, the operating system and the phone together. That lets them add a dedicated block of silicon for this work. It is called the Neural Engine, and it does nothing except the arithmetic models need.

The Neural Engine is fast and very economical with battery. Reaching it needs a format the operating system understands, and that format is Core ML.

What Core ML does that is genuinely clever

You hand over one model file. The system decides, part by part, where each piece should run.

   your .mlpackage file
            |
     [ Core ML decides ]
       /      |      \
    CPU     GPU    Neural Engine
   (odd    (big     (the fast,
    bits)   maths)   cheap one)

You do not choose. The system looks at what the model needs and what the device has. Then it splits the work. On a newer iPhone most of it lands on the Neural Engine.

This is a real convenience, and also the main frustration. When it puts your model somewhere slow, you cannot order it elsewhere.

Getting a model into Core ML

You do not train in Core ML. You train in PyTorch or TensorFlow as usual, then convert with a tool called coremltools.

  PyTorch model  ──►  coremltools  ──►  MyModel.mlpackage  ──►  Xcode  ──►  the app
   (your laptop)                                              (a Mac)

Xcode, Apple's app-building tool, then generates code so the model can be called like a normal function in the app.

The honest part

You need a Mac. Converting the modern format needs macOS or Linux, and running a Core ML model to check it needs macOS specifically. On Windows the conversion of the current format fails outright, and prediction is not available at all.

That is a genuine barrier if you do not own Apple hardware, and most readers of this page do not. Nothing in this lesson pretends otherwise.

Core ML is a dead end for portability. Ship Core ML and you have shipped to Apple. Android needs a separate file in a separate format. Many teams keep ONNX as the single source and convert to each platform from there.

Remember this

  • Core ML is Apple only, and it reaches the Neural Engine.
  • You convert into it; you do not train in it.
  • You need a Mac to convert and test, which is the real barrier for most people.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install coremltools torch scikit-learn

coremltools installs on Windows, macOS and Linux. What it can do differs by platform, and the first script measures exactly that rather than assuming it.

Step 1: convert, and find out what your machine can do

to_coreml.py
import os, platform, numpy as np, torch, torch.nn as nn, torch.nn.functional as F
import coremltools as ct
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

torch.manual_seed(0); torch.set_num_threads(1)
X, y = load_digits(return_X_y=True)
Xtr, Xte, ytr, yte = train_test_split(X / 16.0, y, test_size=0.3, random_state=0, stratify=y)
Xtr = torch.tensor(Xtr, dtype=torch.float32); ytr = torch.tensor(ytr, dtype=torch.long)
Xte = torch.tensor(Xte, dtype=torch.float32)

model = nn.Sequential(nn.Linear(64, 256), nn.ReLU(), nn.Linear(256, 256), nn.ReLU(), nn.Linear(256, 10))
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
for _ in range(30):
    for i in torch.randperm(len(Xtr)).split(64):
        opt.zero_grad(); F.cross_entropy(model(Xtr[i]), ytr[i]).backward(); opt.step()
model.eval()

traced = torch.jit.trace(model, torch.zeros(1, 64))   # coremltools converts a traced graph
print("this machine :", platform.system())

def size_of(path):                                    # .mlpackage is a folder, .mlmodel is a file
    if os.path.isfile(path):
        return os.path.getsize(path)
    return sum(os.path.getsize(os.path.join(r, f)) for r, _, fs in os.walk(path) for f in fs)

def try_convert(fmt, path):
    try:
        m = ct.convert(traced, convert_to=fmt, inputs=[ct.TensorType(name="pixels", shape=(1, 64))])
        m.save(path)
        print("%-14s -> %-18s %.1f KB" % (fmt, path, size_of(path) / 1024))
        return m
    except Exception as e:
        print("%-14s -> failed: %s: %s" % (fmt, type(e).__name__, str(e).splitlines()[0][:60]))
        return None

nn_model = try_convert("neuralnetwork", "digits.mlmodel")     # the older format
ml_model = try_convert("mlprogram", "digits.mlpackage")       # the current one

target = ml_model or nn_model
try:
    out = target.predict({"pixels": Xte[:1].numpy()})
    print("prediction   :", int(np.argmax(list(out.values())[0])))
except Exception as e:
    print("prediction   : not available here -", type(e).__name__)
    print("               ", str(e).splitlines()[0][:80])
Output
this machine : Windows
neuralnetwork  -> digits.mlmodel     332.5 KB
mlprogram      -> failed: RuntimeError: BlobWriter not loaded
prediction   : not available here - Exception
                Model prediction is only supported on macOS version 10.13 or later.

That is the honest output on a Windows laptop, and coremltools prints several version warnings to stderr alongside it. On macOS both conversions succeed and the prediction line prints a digit.

What the two failures actually mean

BlobWriter not loaded is not a bug in your code. The current mlprogram format stores weights in a separate binary blob written by a native library that ships only for macOS and Linux. Windows gets the Python parts of coremltools and not that library.

Model prediction is only supported on macOS is more fundamental. Core ML is the operating system's inference engine. There is no portable implementation to fall back on. On any other platform you can produce a Core ML file and you cannot execute it.

So the workflow if you do not own a Mac: convert to ONNX, verify the model's answers there, and hand the ONNX file to someone with a Mac for the final conversion and testing. Do not ship a Core ML model that nobody has run.

Step 2: shrink it

Core ML models are large by default, because the default is float32. This runs anywhere neuralnetwork conversion worked.

shrink.py
import os, coremltools as ct
from coremltools.models.neural_network.quantization_utils import quantize_weights

base = ct.models.MLModel("digits.mlmodel")
print("float32        %6.1f KB" % (os.path.getsize("digits.mlmodel") / 1024))
for bits in [8, 4]:
    path = "digits_%dbit.mlmodel" % bits
    quantize_weights(base, nbits=bits).save(path)
    print("%2d-bit weights %6.1f KB" % (bits, os.path.getsize(path) / 1024))
Output
float32         332.5 KB
Quantizing using linear quantization
Optimizing Neural Network before Quantization:
Finished optimizing network. Quantizing neural network..
Quantizing layer linear_0 of type innerProduct
Quantizing layer linear_1 of type innerProduct
Quantizing layer linear_2 of type innerProduct
 8-bit weights   87.7 KB
Quantizing using linear quantization
Optimizing Neural Network before Quantization:
Finished optimizing network. Quantizing neural network..
Quantizing layer linear_0 of type innerProduct
Quantizing layer linear_1 of type innerProduct
Quantizing layer linear_2 of type innerProduct
 4-bit weights   46.2 KB

332.5 KB to 87.7 KB at 8 bits, and 46.2 KB at 4 bits. The same 3.8x and 7.2x you saw with every other toolkit, because it is the same arithmetic underneath.

Notice what is missing from that output: any accuracy number. This machine cannot run a Core ML model, so it cannot tell you what the quantisation cost. Never ship a compressed model whose accuracy you have not measured on the platform that will run it.

The modern compression API, for Mac users

On macOS with mlprogram models, the current API is coremltools.optimize.coreml, and it is more capable than the function above:

python
import coremltools as ct
import coremltools.optimize.coreml as cto

model = ct.models.MLModel("digits.mlpackage")

linear = cto.linear_quantize_weights(model, cto.OptimizationConfig(
    global_config=cto.OpLinearQuantizerConfig(mode="linear_symmetric", dtype="int8")))
linear.save("digits_int8.mlpackage")

palettised = cto.palettize_weights(model, cto.OptimizationConfig(
    global_config=cto.OpPalettizerConfig(nbits=4, mode="kmeans")))
palettised.save("digits_4bit.mlpackage")

No output block for this one, deliberately. It needs a Mac, and inventing sizes and accuracies for a machine I did not run it on would teach you to expect numbers that will not appear.

Palettisation is worth understanding because it is Apple's preferred method and it differs from ordinary quantisation. Instead of rounding weights onto an evenly spaced grid, it clusters them with k-means and stores a cluster index per weight plus a small lookup table. At 4 bits that is 16 cluster centres, placed where the weights actually are rather than spread evenly. For weight distributions with a sharp peak at zero — which is most of them — it loses less than linear quantisation at the same bit width.

Common mistakes

Converting a model in training mode. Trace after model.eval(), or dropout and batch-norm behaviour gets baked in wrong.

Assuming the Neural Engine ran your model. Core ML decides silently. Use the Performance report in Xcode 14 and later, which shows every layer and which unit executed it. Teams routinely discover the whole model ran on the CPU.

Fixed input shapes. A model traced at (1, 64) accepts exactly that. Use ct.RangeDim for a flexible dimension, and be aware that flexible shapes reduce how much Core ML can optimise ahead of time.

Forgetting image preprocessing. ct.ImageType lets you fold scaling and mean subtraction into the model, so the Swift side hands over a CVPixelBuffer and nothing else. Doing that arithmetic in Swift instead is a common source of silently wrong results.

Testing only on the simulator. The iOS Simulator has no Neural Engine. Simulator timings tell you nothing about device timings.

Try it yourself

Run step 1 and note the exact failure lines for your platform. Then export the same model to ONNX using the script in the ONNX lesson and compare the file sizes: 332.7 KB for ONNX against 332.5 KB for Core ML. Two formats, the same weights, no compression from either. The format is a container, and the compression is a separate decision.

What to learn next

Researcher — Mathematics and papers.

Two model formats under one name

Core ML has carried two distinct internal representations, and the distinction determines which tools apply.

Neural Network (.mlmodel, 2017) is a protobuf describing a fixed graph of layer types. Weights live inline. It is the format that converts on any platform, and it is in maintenance.

ML Program (.mlpackage, Core ML 5 / iOS 15, 2021) is a directory containing a MIL — Model Intermediate Language — program plus a separate weight blob. MIL is a typed SSA intermediate representation with functions, blocks and explicitly typed operations. It supports typed float16 execution as a first-class choice, stateful models with mutable buffers, and constant expressions that are materialised at load time rather than stored expanded.

The weight blob is why Windows conversion fails: BlobWriter is a compiled component distributed only for macOS and Linux.

The KV-cache support added for ML Program in iOS 18 is what made on-device transformer decoding practical, because it removed the need to re-materialise the cache through the graph on every token.

The compute units

MLComputeUnits gives four settings: .all, .cpuAndGPU, .cpuAndNeuralEngine, .cpuOnly. These are hints. The runtime partitions the graph and assigns segments, and there is no supported API to force a specific unit for a specific layer.

Neural Engine constraints that determine whether a layer is eligible, and which are documented mainly through Apple's Deploying Transformers on the Apple Neural Engine article (2022):

  • Computation is float16. Weights may be stored at lower precision and are expanded on load.
  • The preferred tensor layout is (B, C, 1, S) — batch, channels, one, sequence — rather than the (B, S, C) that transformer implementations naturally produce. Reshaping to match this is the single largest lever in that article's reported 10x improvement.
  • Some operations are unsupported and force a segment onto the CPU or GPU. Each hand-off costs a synchronisation.

The practical consequence: an architecture that is faster in FLOP terms can be slower on Apple hardware if it fragments the graph. Measure with the Xcode Performance report, which shows per-layer unit assignment, before rewriting anything.

Palettisation, stated properly

Weight palettisation replaces each weight with an index into a learned codebook. For $n$ bits and a group of $N$ weights, the storage is

$$ N n + 2^{n} \cdot 16 \ \text{bits} $$

where the second term is the codebook, stored at float16, and the codebook is chosen by k-means over the weight values in the group.

Unlike affine quantisation, the reconstruction levels are not evenly spaced. Since trained weight distributions are approximately Gaussian and sharply peaked at zero, placing levels at quantiles rather than uniformly reduces expected squared error at equal bit width. This is the same insight behind the NF4 data type in QLoRA (Dettmers et al., 2023).

Core ML supports per-grouped-channel codebooks, so group_size trades codebook storage against fidelity in the same way as group-wise affine quantisation. iOS 18 added learned palettisation through the coremltools.optimize.torch path, which fine-tunes with the codebook in the loop rather than fitting it post hoc.

Where Core ML sits against the alternatives

Core MLLiteRTONNX Runtime
PlatformsApple onlyAndroid, iOS, Linux, microNearly everything
Reaches the Neural EngineYesOnly via Core ML delegateOnly via Core ML provider
FormatProtobuf or MIL packageFlatBufferProtobuf
Conversion hostmacOS or LinuxAnyAny
Execution off-platformNoneYesYes

Both LiteRT and ONNX Runtime reach the Neural Engine by delegating to Core ML underneath, which means they inherit Core ML's partitioning behaviour plus one more layer of possible fallback. Direct Core ML conversion remains the route with the fewest places to silently lose acceleration.

Reading

What to learn next